Repository navigation
feat: Only filter records that directly match a kept record - #89
Merged
Merged
Conversation
Pringled
marked this pull request as draft
September 25, 2026 14:28
…eighbor limit Neighbors propose their canonical (in both directions) and each candidate is verified by direct cosine similarity, so matches stay non-transitive. When a record's neighbor list is truncated at MAX_NEIGHBORS without a match, it is compared against all selected records directly.
Codecov Report✅ All modified and coverable lines are covered by tests.
🚀 New features to boost your workflow:
|
Zero embeddings (e.g. empty text) produced NaN normalized vectors, which could make the all-selected fallback pick NaN and keep an extra canonical.
Cross-dataset deduplication now verifies ANN neighbors with exact cosine similarity, like self-deduplication, so scores are exact and zero vectors (e.g. empty text) are never reported as near duplicates.
Pringled
marked this pull request as ready for review
September 26, 2026 11:26
… deprecating duplicates
stephantul
approved these changes
Sep 27, 2026
Pringled
added a commit
that referenced
this pull request
Sep 28, 2026
…h benchmarks and bump to 0.5.0 Representative selection encoded its top candidates again after ranking. The ranking now returns the embeddings in ranked order (index vectors in self mode, the query embeddings in cross mode) and diversification reuses them. Benchmarks are rerun on a MacBook Pro (Apple M5, 48 GB RAM) to reflect #89 and #90, and the version is bumped to 0.5.0 for their breaking changes.
Pringled
added a commit
that referenced
this pull request
Sep 30, 2026
…h benchmarks and bump to 0.5.0 Representative selection encoded its top candidates again after ranking. The ranking now returns the embeddings in ranked order (index vectors in self mode, the query embeddings in cross mode) and diversification reuses them. Benchmarks are rerun on a MacBook Pro (Apple M5, 48 GB RAM) to reflect #89 and #90, and the version is bumped to 0.5.0 for their breaking changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR makes exact and near duplicates follow the same contract: each filtered record references one retained canonical, and selected_with_duplicates provides the complete groups. To make this work, it also reworks the way we consider duplicates by changing to a "direct" deduplication approach. To illustrate the change here's an example:
Suppose we have a dataset with records A, B, C, D, E and threshold=0.9 where the similarities can be represented like this:
In other words, the following pairs are duplicates:
In this case, there are two "clean" ways to deduplicate:
What we on main is a bit different: we keep A and E, and filter B, C, and D. The logic is this (when going through the records in the example) :
seen = {A, B}(since B is a direct duplicate of A)seen(since dropped records never add their neighbours) → keptSo it's kind of an inbetween of direct and transitive deduplication where we follow the chain up to two steps, so a record is dropped if it is a neighbour of a kept record, or a neighbour of one of that kept record's neighbours.
I've also changed the API a bit. Instead of giving
duplicatesin the format of [(duplicate, score)...] for everyfiltered, we now just give aduplicate_ofandscore, see example below. Also, you can still just get the duplicates of any selected item viaselected_with_duplicates, the format change is infilteredsince each filtered record now lists only the selected record it was removed for, instead of all its neighbours, e.g.:Before:
After:
Anyway, I think doing direct is both simpler and more honest. This is the least "aggressive" way of filtering, but it also has the nicest guarantees for the final dataset. What I don't like about the current logic is that some records get filtered from the final dataset, but if you then check them against the finalized dataset they are actually not a duplicate of anything in the final dataset. On the benchmarks, this change leads to 1-3% less dropped records on average. I'll rerun them in a followup since I plan on making a few more changes in followup PRs.